Papers with Black-box Identity Verification
TRAP: Targeted Random Adversarial Prompt Honeypot for Black-Box Identification (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Model (LLM) services and models often come with legal rules on who can use them and how they must use them. |
| Approach: | They propose a method that uses adversarial suffixes to get an answer from a target LLM. |
| Outcome: | The proposed method detects the LLMs with over 95% true positive rate at under 0.2% false positive rate even after a single interaction. |